Back

Genome Biology

Springer Science and Business Media LLC

Preprints posted in the last 30 days, ranked by how well they match Genome Biology's content profile, based on 637 papers previously published here. The average preprint has a 0.47% match score for this journal, so anything above that is already an above-average fit.

1
AlphaGenome deletion responses complement supervised enhancer-gene relation prediction in primary human astrocytes

Huang, Z.; Huang, R.; Han, J.

2026-08-21 genomics 10.64898/2026.08.14.744988 medRxiv
Top 0.1%
30.5%
Show abstract

Sequence-to-function models predict molecular readouts directly from DNA, but recognizing a functional regulatory element is not equivalent to assigning the gene it regulates. We evaluated whether AlphaGenome deletion responses identify experimentally supported enhancer-gene relations, using a frozen K562 analysis and a primary-human-astrocyte CRISPR interference (CRISPRi) resource external to our analysis. The K562 mean contrast was positive but heavy-tailed, and exact joins showed direct collision with released Gasperini and ENCODE-rE2G resources; we therefore treated K562 as supporting evidence. In a frozen evaluation of 2,307 AstroREG relations, AlphaGenome deletion strength discriminated 133 functional relations from 2,174 well-powered nonfunctional relations (average precision 0.479, enhancer-cluster 95% confidence interval 0.394-0.561, prevalence 0.058; area under the receiver-operating-characteristic curve 0.726, 0.659-0.786). Adding AlphaGenome to distance, ABC score, enhancer length, measured expression and assay-depth context increased enhancer-grouped out-of-fold average precision from 0.396 to 0.534 and improved log loss from 0.169 to 0.150. The authors cross-fitted EGrf score was stronger alone (average precision 0.559); in a post-hoc calibration that held out both gene and enhancer folds, adding AlphaGenome increased average precision from 0.550 to 0.619 (paired enhancer-cluster increment 0.068, interval 0.023-0.115) and improved log loss from 0.143 to 0.132. This comparison had asymmetric inputs: EGrf was supervised on AstroREG labels and used local epigenomic and context features, whereas the AlphaGenome score was not fitted in this study to those labels or that feature panel but was read from a pre-existing primary-astrocyte RNA-seq output track. A post-hoc same-enhancer analysis gave conditional AUC 0.741 (0.663-0.814); a smaller same-gene analysis (34 genes, 155 relations) gave 0.701 (0.571-0.823). AstroREG labels and EGrf outputs were public before AlphaGenomes public release, so this evaluation is external to our study but not a post-release or proven-unseen benchmark. The results support complementary relation-level utility, not EGrf superiority, sequence-only deployment, causal assignment at arbitrary loci or equivalence between sequence deletion and CRISPRi.

2
A comprehensive benchmark of transcriptome-wide fusion detection using long-read RNA sequencing

Dorney, R.; Wu, S.; Hung, J. Y.-H.; Hebbard, L.; Schmitz, U.

2026-08-12 bioinformatics 10.64898/2026.08.07.743439 medRxiv
Top 0.1%
26.5%
Show abstract

Fusion transcripts contribute to cancer, inherited diseases, developmental disorders, and evolution. Long-read RNA sequencing enables direct sequencing of full-length transcripts, creating new opportunities to detect complex fusion architectures, including previously inaccessible multi-segmented fusion transcripts. However, accurate transcriptome-wide fusion detection remains challenging because existing methods struggle to distinguish genuine fusion events from technical artefacts. Here, we present a comprehensive benchmark of transcriptome-wide fusion detection using simulated datasets and transcriptomes from three cancer cell lines across Oxford Nanopore Technologies (ONT) cDNA, PCR-cDNA, and direct RNA sequencing, Pacific Biosciences (PacBio) Kinnex sequencing, Illumina short-read RNA sequencing, six long-read fusion callers, and multiple analysis strategies. False-positive fusion calls remained the dominant limitation across sequencing platforms and algorithms. Increasing sequencing depth improved recall but also amplified spurious fusion calls, whereas higher read-support thresholds improved precision at the expense of sensitivity. ONT PCR-cDNA sequencing combined with CTAT-LR-Fusion achieved the best overall balance between precision and recall, whereas JAFFAL was the only caller to reliably identify simulated tri-gene fusions. Consensus calling reduced false positives but markedly reduced sensitivity, with only one of 400 simulated fusions detected by all six callers. Breakpoint localisation emerged as a major limitation across all methods. Long-read sequencing consistently recovered more validated fusion transcripts than short-read sequencing, enabled detection of complex tri-gene fusions, and produced more biologically plausible fusion landscapes with fewer promiscuous gene partners. Collectively, our results establish the first comprehensive benchmarking framework for transcriptome-wide fusion detection, using long-read RNA sequencing, and provide practical guidance for selecting sequencing workflows and computational strategies, while identifying key priorities for future algorithm development.

3
Evaluating the performance of splicing predictors on thousands of synthetic gene variants

Bellido Molias, F.; Kudla, G.

2026-08-27 genomics 10.64898/2026.08.24.746734 medRxiv
Top 0.1%
26.3%
Show abstract

Computational predictors of RNA splicing are increasingly used to interpret genetic variants and to design synthetic genes, yet they are almost always benchmarked on endogenous human sequences closely related to their training data. Whether their performance reflects genuine recognition of splicing signals, or instead exploits statistical features of natural genomes such as conservation and exon-intron composition, remains unclear. Here we benchmark eleven splicing predictors on thousands of synthetic GFP variants that are heavily recoded and dissimilar from any training data, using long-read sequencing to measure splicing directly at each position. Despite this distribution shift, modern deep-learning predictors retained strong performance, and the resulting ranking was largely stable across position-level and construct-level benchmarks. SpliceTransformer ranked highest, followed by AlphaGenome and SpliceAI. Tools that ignore long-range sequence context performed substantially worse, largely because they assign high scores to many non-spliced positions. This ranking broadly agrees with benchmarks on endogenous variants, indicating that the leading models capture transferable, sequence-intrinsic determinants of splicing. We further provide a unified calibration that maps each predictor's scores onto the measured fraction of spliced reads, allowing scores to be interpreted as splicing outcomes and compared directly between tools. Our results show that current deep-learning models generalise beyond natural genomes and provide a practical framework for splicing-aware sequence design.

4
Nonparametric kernel-based detection of spatially variable genes with adaptive shrinkage and scalable multi-sample inference

Ghosh, T.; Ghosh, D.

2026-08-18 bioinformatics 10.64898/2026.08.10.744047 medRxiv
Top 0.1%
26.2%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWIdentifying spatially variable genes (SVGs), genes whose expression varies coherently across tissue space, is a central analytic goal in spatially resolved transcriptomics. Current methods rank spatially variable genes using either significance probabilities from parametric models or effect sizes such as the proportion of spatial variance, but parametric approaches impose distributional assumptions, such as Gaussian processes or negative binomial models, that may be violated for sparse or zero-inflated data. Furthermore, most detection tools cannot jointly model multiple biological replicates, and no existing framework provides both nonparametric significance probabilities and stabilized effect-size estimates with formal uncertainty quantification. Here, we introduce CytoKspace, a nonparametric framework that combines a sparse exponential kernel constructed from nearest-neighbor graphs with a quadratic-form test statistic and adaptive permutation testing. CytoKspace employs a multi-stage adaptive permutation schedule that yields substantial computational savings over fixed-permutation baselines, an adaptive shrinkage layer built on empirical Bayes estimation that stabilizes raw spatial effect sizes and provides posterior estimates with local false sign rates, and a scalable multi-sample extension via Fisher combination of significance probabilities and inverse-variance-weighted meta-analysis that accommodates studies with multiple biological replicates. In extensive simulations across a broad range of sample sizes, gene counts, spatially variable gene fractions, and effect sizes, as well as in applications to two real datasets from the Visium and seqFISH platforms, CytoKspace demonstrates competitive sensitivity, well-calibrated false positive rates, and practical computational requirements compared to existing methods. A software implementation of our method is freely available at https://github.com/Ghoshlab/CytoKspace.

5
Contextual Evaluation of MicroRNA Sequencing Data Harmonization: Performance in Sample Clustering

Zou, J.; Duren, Y.; Wang, X.; Xiang, Y.; Qi, Y.; Wang, M.; Wu, Y.; Singer, S.; Qin, L.-X.

2026-08-14 bioinformatics 10.64898/2026.08.09.743718 medRxiv
Top 0.2%
21.8%
Show abstract

Reliable translation of microRNA sequencing data depends on effective harmonization to mitigate artifacts from variable experimental handling. Although many harmonization methods exist, prior evaluations have focused mainly on differential expression analysis, leaving the impact on subgroup discovery understudied. We present a framework for evaluating harmonization in the context of sample clustering, which integrates AI-augmented datasets, statistical evaluation pipelines, and accessible software tools, enabling systematic comparisons across diverse signal-to-artifact ratios and cluster composition settings. Using this framework, we show that harmonization can, often partially, restore clustering accuracy lost to artifacts, especially at moderate signal-to-artifact ratios, with the level of gains depending on the specific harmonization method, the paired clustering technique, and the cluster composition setting. We further confirm these findings by analyzing reconstructed cohorts from The Cancer Genome Atlas breast cancer microRNA sequencing data. Collectively, the results underscore the need for tailored harmonization to support reliable subgroup discovery and highlight the broader importance of context-specific workflows in translational genomics.

6
FP8 Inference in Genomic Foundation Models: Theoretical vs. Realized Speedups on GenomeOcean

Yu, M.; Egan, R.; Liu, F.; Wang, Z.; Shi, L.

2026-08-14 bioinformatics 10.64898/2026.08.09.743676 medRxiv
Top 0.3%
18.8%
Show abstract

Genomic Foundation Models (GFMs) are increasingly used for large-scale sequence analysis and generation. Compared with frontier language models, GFMs are typically smaller and frequently operate on long genomic sequences, with evaluation often requiring preservation of biologically meaningful structure and sequence-level relationships. Although low-precision post-training quantization (PTQ) has shown substantial memory and throughput benefits for general-purpose language models, it remains unclear whether these benefits transfer to GFMs given their distinct model scales, sequence characteristics, and evaluation requirements. We present an empirical case study of FP8 post-training quantization applied to GenomeOcean, a computationally efficient genomic foundation model with strong reported performance across diverse genomics tasks [Zhou et al., 2025]. Its range of model scales, from 100M to 4B parameters, provides a useful setting for examining how quantization effects vary with model size. We evaluate FP8 across two primary GFM inference regimes--embedding extraction and autoregressive generation--and assess its impact along two dimensions: biological fidelity relative to BF16 baselines and system-level efficiency in terms of throughput, memory usage, and energy efficiency. We find that FP8 largely preserves biological fidelity across the evaluated scales and inference regimes, while reducing GPU memory footprint at 4B scale and improving energy efficiency during autoregressive generation. However, realized throughput gains remain substantially below FP8s theoretical 2x hardware ceiling, with a best-case improvement of 19.3% in autoregressive generation and benefits varying strongly by model scale and workload. Autoregressive generation shows the clearest gains, driven largely by KV-cache compression, whereas embedding extraction provides limited or negative throughput benefits at smaller model scales. We attribute this theory-practice gap to the interaction of model-scale effects, memory-system bottlenecks, and software-stack limitations. These findings highlight the need for workload-specific empirical evaluation before adopting low-precision inference in scientific foundation models. Code availabilityhttps://github.com/jgi-genomeocean/genomeocean_efficiency

7
nf_xpatial: A Reproducible Framework for Standardized Preprocessing and Clustering of Xenium Data

Potter, L. A.; Trull, A.; Kumar, N.; Drake, O. R.; Nogueira, M.; Peters, J.; Heinsbroek, J. A.; Day, J. J.; Worthey, E. A.; Ianov, L.

2026-08-29 bioinformatics 10.64898/2026.08.25.747147 medRxiv
Top 0.3%
18.7%
Show abstract

Recent advances in spatial transcriptomics have enabled the profiling of increasingly larger numbers of genes while retaining single-cell and subcellular resolution in situ. However, standardized bioinformatics workflows for analyzing these datasets have lagged behind, with existing pipelines focusing primarily on image processing and cell segmentation. To address this gap, we present nf_xpatial, a best-practices Nextflow pipeline for the downstream analysis of 10x Genomics Xenium data. The pipeline performs quality control, filtering, log and cell area normalization, multi-sample integration, and both expression-driven and spatially informed clustering across systematic parameter sweeps, allowing users to evaluate and compare clustering resolutions and spatial modeling parameters within a single reproducible run. Overall, nf_xpatial streamlines the processing of Xenium data from platform outputs to integrated single-cell and spatial clustering datasets, providing a standardized starting point from which biologists can fine-tune parameters and proceed to hypothesis-driven spatial analyses.

8
BatchRefiner: fast, significant improvement in batch integration of single-cell embeddings with ensemble refinement

Schäffer, D. E.; Kang, H.; Aksu, E. D.; Edelman, D.; Berger, B.

2026-08-26 bioinformatics 10.64898/2026.08.21.746347 medRxiv
Top 0.3%
18.3%
Show abstract

Data from single-cell RNA sequencing (scRNA-seq) and the Assay for Transposase-Accessible Chromatin (scATAC-seq) are high-dimensional, sparse, and undesirably capture technical variability between experiments or batches. Many analysis methods thus seek to produce a low-dimensional cell-by-feature embedding space that groups together biologically similar cells across batches while distancing dissimilar cells. Here, we introduce ensemble refinement for scRNA-seq and scATAC-seq embeddings, inspired by ensemble methods from statistical machine learning, and implement BatchRefiner, a fast post-processing tool to enhance batch integration. We extensively benchmark widely-used scRNA-seq embedding methods on both batch integration and biological conservation over a wide range of datasets, before and after the addition of BatchRefiner. We extend these benchmarking approaches to provide the first comprehensive benchmark of batch integration for scATAC-seq embedding methods, including BatchRefiner. Importantly, we formalize a significance statistic, which we use to demonstrate BatchRefiner's significant improvement in batch integration across a wide range of embedding methods, atlas-scale datasets, and established metrics.

9
TriTower-m6Am: a triple-tower heterogeneous deep learning architecture integrating semantic, sequential, and structural information for mRNA N6,2'-O-dimethyladenosine site prediction

Xiong, K.; Jia, J.

2026-08-13 bioinformatics 10.64898/2026.08.13.744628 medRxiv
Top 0.3%
18.3%
Show abstract

BackgroundN6,2-O-dimethyladenosine (m6Am) is a cap-proximal mRNA modification deposited by PCIF1 at the first transcribed nucleotide of eukaryotic mRNAs. Knowing where m6Am sites sit across the transcriptome would help explain how cells tune mRNA stability and translation, but current computational predictors typically depend on a single sequence representation and do not jointly model the semantic, sequential, and structural signals carried by an RNA sequence. ResultsWe present TriTower-m6Am, a triple-tower architecture that combines three representations: semantic (RNA-FM with BellPooling), sequential (One-Hot BiLSTM), and structural (RGCN with three typed edges). On an independent test set of 640 sequences, TriTower-m6Am reaches AUC = 0.776, MCC = 0.440, and SN = 0.888, against DTC-m6Ams AUC = 0.765, MCC = 0.411, and SN = 0.800. The 8.8 percentage-point gain in sensitivity means that, for every 100 real m6Am sites, the model recovers roughly 9 additional sites missed by the previous best method. Among the three towers, RGCN alone gives the strongest single signal, and the AUC-weighted ensemble raises sensitivity from the 0.55-0.76 band of the standalone towers to 0.89. Ablating the RGCN edge types shows that backbone connectivity accounts for most of the structural signal. ConclusionsCombining semantic, sequential, and structural views of the same RNA sequence improves m6Am prediction beyond what any single representation achieves. Because each towers contribution to the final prediction is a readable voting weight rather than a hidden parameter, the model is not a black box: a user can read off which tower drove a given prediction and trace it back to the corresponding representation, without running a separate post-hoc explainer. The same design pattern can be transferred to other RNA modification site prediction tasks. Author summaryPredicting where m6Am modifications occur on messenger RNA is important for understanding how cells regulate transcript stability and translation. Existing computational methods typically encode the RNA sequence in a single way, such as k-mer counts or a one-hot code, and treat the model as a black box that emits a prediction without explaining which features drove it. We built TriTower-m6Am to address both limitations. Our model combines three independent encoders--a pretrained RNA language model for semantic patterns, a bidirectional LSTM for local nucleotide order, and a relational graph convolutional network for the structural fold--and fuses their outputs by AUC-weighted voting, so the contribution of each tower to a given prediction is a readable number rather than a hidden parameter. On an independent benchmark the ensemble improves sensitivity by 8.8 percentage points over the prior best method, and ablating the graphs edge types reveals that linear backbone connectivity, rather than long-range base-pairing, carries most of the structural signal. The same triple-tower pattern can be transferred to other RNA modification site prediction tasks.

10
stCNASim: Allele-aware spatial RNA-seq simulator enables systematic benchmarking of copy number inference

Huang, X.; Huang, R.; Qiao, J.; Huang, Y.

2026-08-11 bioinformatics 10.64898/2026.08.06.743179 medRxiv
Top 0.4%
18.2%
Show abstract

Spatial transcriptomics (ST) is revolutionizing the study of tumor evolution by enabling spatially resolved copy-number alteration (CNA) analysis. However, evaluating the accuracy and robustness of current single-cell (SC) and ST-specific CNA inference tools remains challenging due to the absence of ground-truth datasets. Here, we present stCNASim, an allele-aware spatial RNA-seq simulator that generates raw reads within realistic spatial contexts. We synthesized 46 benchmarking datasets across varying technical settings and spatial architectures to evaluate five widely used computational methods. Our analysis reveals that while SC-based methods adapt well to ST data, ST-specific methods successfully benefit from considering spatial autocorrelation but struggle under high spatial intermixing. The allele-aware methods CalicoST, Numbat, and XClone achieved top-tier performance with unique advantages in extreme scenarios, yet showed distinct sensitivities to low purity, mirrored alleles, and low coverage, respectively. By providing a scalable simulator and a rigorous benchmark, this work establishes a much-needed framework to guide and accelerate future tool development in spatial CNA analysis.

11
Automated generation of a gene perturbation transcriptomic atlas using large language models

Soul, J.; Young, D. A.

2026-08-14 bioinformatics 10.64898/2026.08.08.743502 medRxiv
Top 0.4%
18.1%
Show abstract

Public transcriptomic repositories contain thousands of gene perturbation experiments, a valuable resource for understanding gene function, but perturbation metadata are not structured, which blocks systematic reuse. Existing perturbation atlases depend on expert manual curation, so they are costly to maintain and infrequently updated, while automated grouping approaches neither identify which samples form the perturbation arm nor recover the perturbed gene. Here we develop an automated pipeline that uses large language models to find single-gene perturbation experiments in NCBI-GEO and reconstruct their case-control sample groupings, along with the perturbed gene, perturbation type and cell line as structured, ontology-normalised fields. We manually curated 3,300 GEO experiments with sample-level case-control assignments and release these as an open benchmark (2,400 training, 600 validation, 300 temporally held-out test). Reasoning models and task-specific finetuning substantially improved identification of valid perturbation groups, with the best model reaching precision 0.925 and recall 0.836 on the test set. Applied at scale, the pipeline generated an atlas of 6,802 gene perturbation expression signatures from 4,453 GEO experiments, covering 2,907 uniquely perturbed genes. An R package, perturbMatch, supports exploration of the atlas and querying of user-supplied expression signatures against it using similarity scoring, so users can identify experiments that recapitulate a transcriptional state of interest.

12
Trust-Aware Sequence-to-Function Modelling in Regulatory Genomics

Onawole, A.; Basiru, S.; Sanni, M. O.; Aiyedun, M.; Sulaimon, R.

2026-08-24 genomics 10.64898/2026.08.20.745945 medRxiv
Top 0.4%
17.9%
Show abstract

Objective: Sequence-to-function models increasingly predict regulatory activity, such as chromatin accessibility, directly from DNA sequence, and are used to interpret non-coding genetic variation. Standard accuracy metrics, computed over a held-out set of genomic regions, do not establish whether an individual prediction remains reliable once the input sequence departs from that set, nor whether a model's attribution-based explanation is biologically grounded rather than coincidental. We develop and evaluate RegTrust-XAI, a trust-aware framework separating these questions using three inference-time signals: ensemble consensus, motif-grounded attribution coherence, and applicability-domain distance. Methods: A five-model convolutional ensemble was trained on 517,790 K562 ATAC-seq windows and evaluated on a held-out chromosome test set (chr8/chr9, n = 42,844). Consensus, coherence, and applicability-domain distance were each tested against prediction error, alongside complementary sequence-novelty analyses and validation against an independent lentiMPRA reporter assay and saturation-mutagenesis MPRA data at the PKLR promoter. Results: The ensemble reached Spearman {rho} = 0.782, with skill of 0.328 over a constant-value null predictor. High-consensus predictions (Scenarios A+B) were consistently enriched for lower error than low-consensus predictions (Scenarios C+D), and attribution coherence further separated error within the high-consensus population (mean absolute error 0.396 versus 0.435, p = 9.6e-10). Applicability-domain distance showed a monotonic error gradient across six distance bands. A 4-mer composition-divergence metric was negatively associated with error and anti-correlated with applicability-domain distance, so composition-based and model-relevant novelty are not equivalent. Attribution transfer to lentiMPRA was assay- and subgroup-dependent, and predicted allele-substitution effects correlated with measured saturation-mutagenesis effects at the PKLR promoter at both 24 h and 48 h ({rho} = 0.227 and 0.235). Motif-specific perturbation further showed that regulatory attributions were strongly context-dependent, with more than 90% of multi-instance motif modules exhibiting superadditive joint effects. Conclusions: Prediction reliability, explanation validity, and sequence novelty are related but distinct properties of a sequence-to-function model. Evaluating each explicitly gives a more complete basis for deciding when to act on a prediction than accuracy alone.

13
SuSiNE: Genetic fine-mapping with signed functional priors and multi-basin ensembling

Callahan, M. G.; Zhu, X.

2026-08-06 genomics 10.64898/2026.07.31.742084 medRxiv
Top 0.5%
15.3%
Show abstract

Genetic fine-mapping identifies causal variants within trait-associated loci, but linkage disequilibrium (LD) and wide datasets complicate this sparse variable-selection problem. SuSiE is popular for its fast variational inference, posterior inclusion probabilities (PIPs), and credible sets, yet a single fit can fail to resolve LD ambiguity, converge to a poor local optimum, or misrepresent uncertainty over competing configurations. We introduce SuSiNE (Sum of Single Non-central Effects), a SuSiE extension incorporating signed functional annotations through a prior-mean channel, {micro}0 = ca, while preserving effect conjugacy, credible sets, and summary-statistic sufficiency. The resulting single-effect Bayes factor self-gates on agreement between annotation sign and association direction, limiting annotation-noise influence. We show that the common final step of purity filtering can discard informative signal, and tends to hurt performance. We also introduce new effect-level diagnostics for concentration, accuracy, and fitted-basis movement, to provide deeper insights into model behavior. To explore and summarize multiple variational basins, we pair the model with grid-based ensembling and cluster-weight aggregation. In oligogenic simulations with annotations calibrated to AlphaGenome eQTL bench-marks, the ensemble raised pooled AUPRC for recovery of the largest-effect causal variants from a SuSiE-equivalent 0.2474 to 0.3130 (0.0656 delta, 95% paired-bootstrap CI [0.0591, 0.0722]). At 75% precision, recall rose from 11.9% to 19.3% (61.7% relative gain). AUPRC gains were robust across varying annotation quality and alternative sparse and diffuse architectures, while sufficiently strong null annotation-association alignment reversed the gains. In a GTEx Lung summary-statistic case study, SuSiNE placed nontrivial weight on annotation-informed fits at 7 of 20 loci and changed which variants received high PIP. ARSA showed the cleanest durable shift, whereas the large YDJC shift coincided with reference-LD discrepancy. An internal diagnostic found little evidence of strong annotation confounding in this panel. These analyses use reference rather than in-cohort LD, demonstrating method behavior rather than definitive variant-level discoveries. Author summaryWhen a genetic study links part of the genome to a disease or to differences in gene expression, the next question is which variants are responsible. Answering this is hard, because nearby variants are usually inherited together and can look almost interchangeable in the data. We studied a widely used method, SuSiE, by asking where it breaks down. We found that a routine final cleanup step often discards real signal for nothing in return. A single run can also settle on one explanation without exploring alternatives that fit the data just as well. We introduce new checks that make both problems visible. We then developed SuSiNE, which lets the method use directional predictions from AI sequence models or other biological evidence. It runs many times across settings that encourage exploration, then combines the results into one summary. In calibrated simulations, SuSiNE found true causal variants substantially more often than the standard method. On real gene-expression data, it changed which variants look responsible at several locations. These results are limited, but they suggest AI sequence models are already good enough to offer competing explanations at well-studied genome locations, if we use them carefully.

14
IsoAtlas: Visual interpretation of known and novel transcript isoforms using population-scale long-read evidence

Zheng, X.; Sedlazeck, F. J.

2026-08-21 bioinformatics 10.64898/2026.08.12.744438 medRxiv
Top 0.5%
15.2%
Show abstract

Long-read RNA sequencing has revealed extensive transcript diversity, but newly observed isoforms remain difficult to interpret beyond their classification as known or novel. Here, we present IsoAtlas, an interactive multispecies database for visual exploration and population-scale interpretation of transcript isoforms using 1, 035 human and 414 mouse uniformly processed long-read RNA-sequencing samples. Users can search annotated genes and transcripts or submit novel transcript models in GTF format, visualize their structures, and examine sample-level support, prevalence, expression, tissue and disease context, and sequencing-platform evidence. IsoAtlas integrates structurally equivalent transcripts across GENCODE, RefSeq and CHESS, consolidating evidence that would otherwise be distributed across annotation-specific identifiers. It further links corresponding human and mouse transcript models, enabling users to assess cross-species conservation and enabling users to assess cross-species conservation and inform the suitability of mouse models for isoform-specific studies. IsoAtlas can also evaluate arbitrary user-supplied transcript structures directly against accumulated long-read evidence. IsoAtlas therefore complements established reference annotations with an extensible evidence layer that connects transcript structure to population prevalence, biological context and cross-species support. IsoAtlas is freely available at https://www.isoatlas.org/. Graphic abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=83 SRC="FIGDIR/small/744438v1_ufig1.gif" ALT="Figure 1"> View larger version (21K): org.highwire.dtl.DTLVardef@89bd72org.highwire.dtl.DTLVardef@f48bbforg.highwire.dtl.DTLVardef@102ae37org.highwire.dtl.DTLVardef@fbbc2b_HPS_FORMAT_FIGEXP M_FIG C_FIG

15
pysigscore: gene signatures scoring across bulk and single-cell transcriptomics

Giacomello, T.; Mazzara, S.; Abbruzzese, G.; Barberis, A.; tangherloni, a.; Buffa, F. M.

2026-08-09 bioinformatics 10.64898/2026.08.04.742537 medRxiv
Top 0.5%
15.1%
Show abstract

SummaryHigh-throughput transcriptomics has made gene signatures central to interpreting gene expression data, with applications in diagnosis, prognosis, and prediction. Quantifying signature activity and assessing its robustness remain challenging because scoring methods primarily rely on various assumptions, and no single approach is universally optimal. Here, we present pysigscore, a Python framework for gene set scoring in bulk and single-cell RNA-seq data. pysigscore integrates 18 built-in scoring methods with a fully customisable scorer, allowing users to define and benchmark new scoring functions. It also provides reliability analyses, including p-value estimation and leave-one-out experiments, to assess the significance of scores and gene-level contributions. We validated pysigscore on the CCLE, TCGA, and PBMC datasets, recovering the expected enrichment in liver, hypoxia, inflammatory, and cell-cycle signatures. Availability and ImplementationSource code is available at https://github.com/bioinformatics-hub/pysigscore. Contact: tommaso.giacomello@phd.unibocconi.it, francesca.buffa@unibocconi.it Supplementary informationSupplementary data are available at Bioinformatics online.

16
Dissecting context-dependent cancer vulnerabilities using Perturb-seq

Maffa, S.; Boyle, I. A.; Ward, L.; Colgan, W. N.; Borck, P.; Simerzin, A.; Adeagbo, A.; Olajide, O.; Wie, S.; Liang, H.; Wienand, K.; Shibue, T.; Ray, J.; Paolella, B.; Campbell, C. D.; Vazquez, F.; Dempster, J. M.

2026-08-25 cancer biology 10.64898/2026.08.24.746802 medRxiv
Top 0.6%
14.8%
Show abstract

Background CRISPR-mediated viability assays in diverse cancer cell lines have informed cancer biology and precision medicine, but cell fitness is not the only cancer-relevant phenotype. Gene expression profiling provides insight into cellular stress, inflammation, and differential state, while still identifying activation of cell-death pathways. Perturb-seq allows scalable functional genomics screening of expression phenotypes at single-cell resolution, however existing datasets cover only a small number of work-horse cell lines. Results We produced a proof-of-concept Perturb-seq dataset targeting 100 genes in 16 diverse cancer cell lines. In the process, we established methods to address single-cell technical artifacts, identified Cas9-mediated chromosomal aberrations and assessed screen quality. Even with a limited library, we observed common signatures of deleting essential genes as well as context-specific responses based on intrinsic genomic properties of the models. For example, we inferred a previously undescribed relationship between dependence on the ER-golgi transport gene immediate early response 3 interacting protein 1 (IER3IP1) and oxidative stress, demonstrating the potential of integrated Perturb-seq for hypothesis generation. Conclusions We established a framework for building a comprehensive map of post-perturbational transcriptional phenotypes using parallel Perturb-seq experiments across multiple cell lines. We demonstrated that integrated Perturb-seq experiments spanning diverse contexts enable hypotheses about gene function specific to tissue types or cancer subtypes - suggesting large-scale, genome-wide datasets would offer invaluable insight into the highly context-dependent nature of cancer biology.

17
Phylogeny-aided detection of contamination in nearly 5 million SARS-CoV-2 genomes

Anoufa, O.; Ly-Trong, N.; Goldman, N.; De Maio, N.

2026-08-07 bioinformatics 10.64898/2026.08.07.743473 medRxiv
Top 0.6%
14.7%
Show abstract

Contamination can occur during genome sequencing when a contaminant genome is accidentally mixed with the intended genome to be sequenced. Contamination can lead to incorrect consensus genome calling, disrupting analyses of pathogen evolution and transmission. To investigate the extent of this issue, we developed PhyCD, a phylogeny-aided computational approach to investigate contamination in SARS-CoV-2 genome sequencing data. PhyCD masks consensus genome positions associated with suspicious sequencing read coverage drops, then leverages pandemic-scale phylogenetic placement techniques to identify putative contamination events. Applying PhyCD to nearly 5 million SARS-CoV-2 genomes, we identified 10,942 putative contamination events under conservative parameters. Across the flagged genomes, PhyCD flagged -- and so permits masking of -- a total of 64,753 substitutions that could cause errors in downstream genome data analyses.

18
Flex-sweep 2.0: more flexible and faster selective sweeps detection

Murga-Moreno, J.; Enard, D.

2026-08-07 evolutionary biology 10.64898/2026.08.06.743046 medRxiv
Top 0.6%
14.5%
Show abstract

Flex-sweep is a convolutional neural network-based method able to detect a wide range of selective sweeps, including those thousands of generations old, from single population genomic data, while robust to background selection. Here we present a substantial update that streamlines the entire workflow. The new version vastly reduces memory needs and vastly speeds up summary-statistic computation over fully customizable statistics combinations and genomic regions, relaxes CNN constraints by supporting custom architectures and haplotype matrix sorting methods. Domain-Adaptive Neural Network (DANN) training is now supported, as well as ancestral-state polarization and a robust, clustering and confounder-aware gene set sweep enrichment pipeline robust for downstream analysis. Flex-sweep 2.0 scales to hundreds of thousands of training simulations, and enables genome-wide inference on a standard workstation.

19
On the robustness of scRNA-seq foundation models for plant perturbation response prediction under cross-experiment shift

Fernandez Burda, M.; Bonazzola, R.; Valli, A. A.; Castrillo, G.; Stegmayer, G.; Ferrante, E.; Milone, D. H.

2026-08-22 bioinformatics 10.64898/2026.08.21.746324 medRxiv
Top 0.6%
13.1%
Show abstract

Foundation models for single-cell transcriptomics promise to learn generalizable representations of cellular states. However, recent evidence suggests they often fail to outperform simple machine learning baselines. Furthermore, their ability to generalize across unseen experimental conditions remains poorly understood, particularly in plants, where rigorous evaluation beyond cell type annotation and batch integration is lacking. To address this, we introduce an Arabidopsis thaliana foundation model, scAraFM, and benchmark it across several perturbation conditions under three increasingly challenging protocols: random splits from a single experiment, replicate-based splits, and cross-experiment transfer learning. We found that random splits overestimate performance by up to 30 points relative to cross-experiment evaluations. Across representation strategies, preserving gene identity consistently outperforms the standard pooled embeddings. Moreover, simple baselines using raw reads remain competitive in single-experiment settings, challenging current claims of universal advantage of foundation models. In contrast, under cross-experiment transfer, pretrained representations show added value, particularly with few labelled samples, suggesting that the benefits of foundation models emerge precisely in the regimes that matter for practical deployment. Overall, our results demonstrate that conclusions about foundation models depend critically on the evaluation design, and that preserving per-gene structure aids generalization in downstream tasks, supporting robust predictions across unseen experimental contexts.

20
Comparative Single-Cell Profiling of CRISPR Knockout and Interference Defines Modality-Specific Strengths in Functional Genomics

Gupta, N.; Sayer, A.; Prater, M.; Mastrokalou, C.; Saeed, K.; Company, C.; Hart, C.; Miragaia, R.; Trehan, A.; Functional Genomics Centre, ; McDermott, U.; Strauss, M. E.; Ross-Thriepland, D.; Walter, D.; Kalinka, A.

2026-08-27 genomics 10.64898/2026.08.26.747303 medRxiv
Top 0.6%
13.1%
Show abstract

Pooled CRISPR screens coupled with single-cell RNA sequencing enable high-throughput functional interrogation of gene regulatory networks, yet systematic comparisons of CRISPR knockout (CRISPRko) and CRISPR interference (CRISPRi) remain limited. We established a single-cell CRISPRko workflow and benchmarked it against CRISPRi using 87 sgRNAs targeting 29 unfolded protein response genes. CRISPRko generated transcriptional phenotypes are highly concordant with CRISPRi, and induced comparable pathway-level responses. While CRISPRi allows direct assessment of target gene repression, CRISPRko provides an effective complementary approach for complete loss-of-function studies, expanding the toolkit for single-cell functional genomics.